Papers with reinforcement learning agent
Deep Reinforcement Learning for Chinese Zero Pronoun Resolution (P18-1)
Copied to clipboard
| Challenge: | Recent models for zero pronoun resolution in Chinese are short-sighted and do not capture semantic information for zeros and candidate antecedents. |
| Approach: | They propose to integrate a deep reinforcement learning approach to Chinese zero pronoun resolution. |
| Outcome: | The proposed approach outperforms the state-of-the-art methods in three experimental settings. |
Enhancing multi-modal Relation Extraction with Reinforcement Learning Guided Graph Diffusion Framework (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for cross-modal relation extraction focus on single-modal data, which limits their use in real-world situations. |
| Approach: | They propose a framework that leverages pre-trained models to encode multi-modal data into scene graphs and combine them into a cross-modal graph. |
| Outcome: | The proposed model outperforms existing methods on multi-modal relation extraction tasks. |
Mapping Smarter, Not Harder: A Test-Time Reinforcement Learning Agent That Improve Without Labels or Model Updates (2025.emnlp-industry)
Copied to clipboard
| Challenge: | a new agent that can improve schema mappings for third-party logs is needed for enterprise intelligence platforms. |
| Approach: | They propose a reinforcement learning agent that can self-improve without labeled examples or model weight updates. |
| Outcome: | The proposed method increases mapping accuracy from 56.4% (LLM-only) to 72.73% (RAG) to 93.94% over 100 iterations using GPT-4o. |
Posterior-regularized REINFORCE for Instance Selection in Distant Supervision (N19-1)
Copied to clipboard
| Challenge: | Existing methods to train unbiased methods such as REINFORCE take time to train. |
| Approach: | They propose to use posterior regularization to integrate domain-specific rules in instance selection using REINFORCE to improve the performance of the relation classifier trained on cleaned distant supervision datasets. |
| Outcome: | The proposed method improves the performance of the relation classifier trained on cleaned distant supervision dataset as well as the efficiency of the REINFORCE training. |
Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee Recognition (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to learn dialogue discourse parsing with related tasks require additional annotation, thus limiting their generality. |
| Approach: | They propose a multitasking framework that integrates dialogue discourse parsing with addressee recognition to reflect relation-based structure of dialogue. |
| Outcome: | The proposed framework outperforms baselines on the Molweni and STAC datasets. |
Keep CALM and Explore: Language Models for Action Generation in Text-based Games (2020.emnlp-main)
Copied to clipboard
| Challenge: | Text-based games present a unique challenge for autonomous agents to operate in natural language and handle enormous action spaces. |
| Approach: | They propose a Contextual Action Language Model (CALM) to generate a compact set of action candidates at each game state. |
| Outcome: | The proposed model achieves a 69% improvement in average game score on unsupervised games . the proposed model is competitive with or better than other models that have access to ground truth admissible actions on half of the games tested . |